Papers with Fast Vocabulary Transfer

2 papers
Efficient Low-Resource Language Models Using Tokenizer Transfer (2026.eacl-srw)

Copied to clipboard

Challenge: Tokenizer transfer allows training a model for low-resource languages without full retraining . a study of pre-trained tokenizers shows that they are more efficient than traditional training methods.
Approach: They evaluate tokenizer transfer on models trained on language-specific corpora, Orthogonal Mapping Pursuit and Fast Vocabulary Transfer.
Outcome: The proposed model adapts to a pre-trained model without full retraining and improves cross-lingual applicability.
Evaluating Tokenizer Adaptation Methods for Large Language Models on Low-Resource Programming Languages (2025.acl-srw)

Copied to clipboard

Challenge: Large language models (LLMs) trained on high-resource programming languages perform sub-optimally for low-resourced programming languages (LRPLs).
Approach: They evaluate the impact of tokenizer adaptation methods on improving code generation for LRPLs.
Outcome: The proposed methods outperform the original models and fine-tuned models in LRPLs, but performance declines in non-target languages like Python after tokenizer adaptation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations